Papers with adversarial robustness
MENLI: Robust Evaluation Metrics from Natural Language Inference (2023.tacl-1)
Copied to clipboard
| Challenge: | Recent proposed BERT-based evaluation metrics for text generation are vulnerable to adversarial attacks, e.g., relating to information correctness. |
| Approach: | They propose to use BERT-based evaluation metrics for text generation to evaluate text for semantic similarity but are vulnerable to adversarial attacks using Natural Language Inference. |
| Outcome: | The proposed metrics outperform existing summarization metrics but perform below SOTA MT metrics on standard benchmarks. |
In and Out-of-Domain Text Adversarial Robustness via Label Smoothing (2023.acl-short)
Copied to clipboard
| Challenge: | Existing studies show that state-of-the-art NLP models are vulnerable to adversarial attacks . label smoothing has been proven effective in a variety of applications and modalities . |
| Approach: | They propose to use label smoothing to improve adversarial robustness in pre-trained models against various popular attacks. |
| Outcome: | The proposed method significantly improves adversarial robustness in pre-trained models against various popular attacks. |
How to Select One Among All ? An Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Knowledge Distillation (KD) is a model compression algorithm that helps transfer knowledge in a large neural network into a smaller one. |
| Approach: | They propose a framework to assess adversarial robustness of multiple KD algorithms. |
| Outcome: | The proposed algorithm achieves state-of-the-art on the GLUE benchmark and out-of domain generalization and adversarial robustness compared to competitive methods. |
Certified Robustness to Word Substitution Attack with Differential Privacy (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent studies have shown that adversarial examples can be easily fooled by DNNs, making the robustness and security of NLP models significantly important. |
| Approach: | They propose a differential privacy-based algorithm to achieve certified robustness against word substitution at- tacks in text classification via differential privacy. |
| Outcome: | The proposed model achieves higher accuracy and more than 30X efficiency improvement over existing defense algorithms. |
Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have driven major advances across domains, yet their massive size hinders deployment in resource-constrained settings. |
| Approach: | They propose to compress large language models to reduce computation and memory consumption while maintaining accuracy. |
| Outcome: | The proposed algorithms preserve training data privacy but weaken the protection of personally identifiable information during conversations. |
Robust Transfer Learning with Pretrained Language Models through Adapters (2021.acl-short)
Copied to clipboard
| Challenge: | Existing approaches to transfer learning with pretrained transformer-based language models are not robust and can be adversarial. |
| Approach: | They propose a simple yet effective adapter-based approach to fine-tune language models on downstream tasks. |
| Outcome: | The proposed approach improves stability and adversarial robustness in transfer learning to various downstream tasks. |
Towards Adversarially Robust Text Classifiers by Learning to Reweight Clean Examples (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing defense methods improve the adversarial robustness by making models adapt to training set augmented with some adversarials. |
| Approach: | They propose to introduce a reweighting mechanism to calibrate the training distribution to obtain robust models. |
| Outcome: | The proposed method minimizes the loss of validation set mixed with clean examples and adversarial ones in an online learning manner. |
Better Robustness by More Coverage: Adversarial and Mixup Data Augmentation for Robust Finetuning (2021.findings-acl)
Copied to clipboard
| Challenge: | Pretrained language models perform poorly under adversarial attacks due to the large search space. |
| Approach: | They propose a method to cover a much larger proportion of the attack search space by adding textual adversarial examples during training. |
| Outcome: | The proposed method covers a much larger proportion of the attack search space. |
Adversarial Robustness of Prompt-based Few-Shot Learning for Natural Language Understanding (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent few-shot learning methods focus on improving downstream task performance, but there is limited understanding of the adversarial robustness of such methods. |
| Approach: | They evaluate prompt-based FSL methods against fully fine-tuned models to better understand the impact of various factors towards robustness. |
| Outcome: | The proposed methods show that they are less robust in the face of adversarial perturbations than fully fine-tuned models. |
Sibylvariant Transformations for Robust Text Classification (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing text transformation techniques are limited in their ability to expand input space . many techniques can artificially expand labeled training sets or test suites, but are class-preserving . |
| Approach: | They propose a concept of sibylvariance to describe transforms that relax the label-preserving constraint and knowably vary the expected class. |
| Outcome: | The proposed transforms can expand input space, but they are limited in their ability to expand . the proposed transform can knowably vary the expected class and lead to more diverse distributions . |
ROSE: Robust Selective Fine-tuning for Pre-trained Language Models (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have highlighted the lack of adversarial robustness in pre-trained models. |
| Approach: | They propose a fine-tuning approach that conducts selective updates when adapting pre-trained models to downstream tasks. |
| Outcome: | The proposed approach improves adversarial robustness on downstream tasks . it eliminates spurious updates, leading to flatter and wider optima than the conventional method . |
Generalized but not Robust? Comparing the Effects of Data Modification Methods on Out-of-Domain Generalization and Adversarial Robustness (2022.findings-acl)
Copied to clipboard
| Challenge: | Data modification has been proposed as an effective solution for generalizing to out-of-domain (OOD) inputs. |
| Approach: | They propose to use data modification to generalize to out-of-domain inputs . they also analyze their adversarial robustness using a synthetic dataset . |
| Outcome: | The proposed data modification strategies improve OOD accuracy and AR, but data filtering hurts OOD on other tasks. |
RoAST: Robustifying Language Models via Adversarial Perturbation with Selective Training (2023.findings-emnlp)
Copied to clipboard
Jaehyung Kim, Yuning Mao, Rui Hou, Hanchao Yu, Davis Liang, Pascale Fung, Qifan Wang, Fuli Feng, Lifu Huang, Madian Khabsa
| Challenge: | Several perspectives of robustness for pre-trained language models have been studied independently, but lacking a unified consideration in multiple perspectives. |
| Approach: | They propose a technique to enhance the multi-perspective robustness of LMs by introducing adversarial perturbation while the model parameters are selectively updated upon their relative importance. |
| Outcome: | The proposed technique improves the robustness of LMs by incorporating four perspectives on model robustness. |
Randomized Smoothing with Masked Inference for Adversarially Robust Text Classifications (2023.acl-long)
Copied to clipboard
| Challenge: | Large-scale pre-trained language models are brittle against specifically crafted adversarial examples, leading to increasing interest in probing the adversariality of NLP systems. |
| Approach: | They propose a two-stage framework that combines randomized smoothing and masked inference to improve the adversarial robustness of NLP systems. |
| Outcome: | The proposed framework improves adversarial robustness by 2 to 3 times over existing state-of-the-art methods on benchmark datasets. |
On Evaluation of Adversarial Perturbations for Sequence-to-Sequence Models (N19-1)
Copied to clipboard
| Challenge: | Existing methods for assessing the robustness of sequence-to-sequence models have been ignored by the literature. |
| Approach: | They propose an evaluation framework for adversarial attacks on seq2seq models that takes the semantic equivalence of the pre- and post-perturbation input into account. |
| Outcome: | The proposed framework breaks the assumption that source perturbations should not result in changes in the expected output, but allows for meaning-preserving perturbations that change the output sequence. |
Exploring Robust Overfitting for Pre-trained Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent literature has revealed their vulnerability to crafted adversarial examples on a wide range of NLP tasks. |
| Approach: | They propose to combine regularization methods with adversarial training to mitigate robust overfitting for pre-trained language models by scaling the model's loss. |
| Outcome: | The proposed methods mitigate robust overfitting upon three top adversarial training methods and further promote adversarially robustness. |
Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning (2026.findings-acl)
Copied to clipboard
| Challenge: | Increasing demand for training data is causing language models to be trained on synthetic data, a new study finds . fine-tuning models on synthetic datasets reduces self-preference bias . |
| Approach: | They investigate the impact of diversity of synthetic data on fine-tuned large language models. |
| Outcome: | The proposed model can mitigate distribution collapse, maintain diversity of output distribution, and reduce self-preference bias. |
Authorship Obfuscation in Multilingual Machine-Generated Text Detection (2024.findings-emnlp)
Copied to clipboard
Dominik Macko, Robert Moro, Adaku Uchendu, Ivan Srba, Jason Lucas, Michiharu Yamashita, Nafis Irtiza Tripto, Dongwon Lee, Jakub Simko, Maria Bielikova
| Challenge: | Recent advances in Language Modeling have birthed Large Language Models (LLMs), which exhibit significant improvements, including the ability to generate texts easily misconstrued as humanwritten. |
| Approach: | They compare authorship obfuscation methods against machine-generated text (MGT) in 11 languages and analyze their performance against 37 well-known AO methods. |
| Outcome: | The proposed methods can cause evasion of detection in all languages, with homoglyph attacks particularly successful. |
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in LLMs (2026.eacl-long)
Copied to clipboard
Yujia Zheng, Tianhao Li, Haotian Huang, Tianyu Zeng, Jingyu Lu, Chuangxin Chu, Yuekai Huang, Ziyou Jiang, Qian Xiong, Yuyao Ge, Mingyang Li
| Challenge: | Existing studies treat prompts as flat text, overlooking their internal structure, and different components within a prompt contribute unequally to robustness. |
| Approach: | They propose a framework that decomposes prompts into functional components and a method that selectively modifies components to expose component-wise vulnerabilities. |
| Outcome: | The proposed framework exposes component-wise vulnerabilities while ensuring linguistic plausibility through perplexity-based filtering. |
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are powerful zero-shot assessors used in real-world situations . however, no study has examined the vulnerability of judge-LLM to adversarial manipulation . |
| Approach: | They propose a simple surrogate attack where a surrogated model is attacked and the learned attack phrase transferred to unknown judge-LLMs. |
| Outcome: | The proposed algorithm shows that judge-LLMs can be significantly more susceptible to adversarial attacks when used for absolute scoring, rather than comparative assessment. |
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks (2025.coling-main)
Copied to clipboard
| Challenge: | Recent advances in natural language processing have highlighted the vulnerability of deep learning models to adversarial attacks. |
| Approach: | They propose a benchmark for textual adversarial defence that evaluates state-of-the-art defence mechanisms across diverse datasets, models, and tasks. |
| Outcome: | The proposed benchmark incorporates a wide range of datasets and evaluates state-of-the-art defence mechanisms. |
Weight Perturbation as Defense against Adversarial Word Substitutions (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existence and pervasiveness of textual adversarial examples have raised serious concerns to security-critical applications. |
| Approach: | They propose to perform weight perturbations in the parameter space rather than the input feature space to improve adversarial robustness of NLP models. |
| Outcome: | The proposed method improves adversarial robustness of models by performing weight perturbations in the parameter space rather than the input feature space. |
MisinfoBench: A Multi-Dimensional Benchmark for Evaluating LLMs’ Resilience to Misinformation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks assess factual accuracy in isolated queries but fail to evaluate LLMs’ resilience to misinformation in interactive settings. |
| Approach: | MisinfoBench is a benchmark designed to assess LLMs’ ability to discern, resist, and reject misinformation. |
| Outcome: | MisinfoBench assesses large language models’ ability to discern, resist, and reject misinformation in interactive settings. |
Model-tuning Via Prompts Makes NLP Models Adversarially Robust (2023.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained models are typically adapted to downstream tasks by appending a randomly initialized multilayer perceptron to their topmost representation layer and fine-tuning the entire model on a downstream task. |
| Approach: | They propose to append a multilayer perceptron to a CLS token and fine-tune the entire model on a downstream task. |
| Outcome: | The proposed model-tuning via prompts outperforms adversarial training-based state-of-art defenses by 3.5% and improves against adversarials by 8% over standard methods. |
Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have shown that adversarial examples can alter models' predicted sentiment due to their sensitivity to specific word choices. |
| Approach: | They propose a regularization technique to improve NLP model robustness by reducing the impact of input perturbations on model outputs. |
| Outcome: | The proposed method outperforms state-of-the-art methods in adversarial defense. |
HypMix: Hyperbolic Interpolative Data Augmentation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for data augmentation involve performing mathematical operations over the raw input samples or their latent states representations, but these operations are performed in the Euclidean space, simplifying these representations and resulting in noisy interpolations. |
| Approach: | They propose a model-, data-, and modality-agnostic interpolative data augmentation technique operating in the hyperbolic space that captures the complex geometry of input and hidden state hierarchies better than its contemporaries. |
| Outcome: | The proposed technique outperforms state-of-the-art methods on benchmark and low resource datasets across speech, text, and vision modalities. |
Don’t Retrain, Just Rewrite: Countering Adversarial Perturbations by Rewriting Text (2023.acl-long)
Copied to clipboard
| Challenge: | ATINTER model can be used to rewrite adversarial inputs to make them non-adversarial . if undefended, model should maintain good task performance and effectively mitigate adversarials . |
| Approach: | They propose a model that intercepts adversarial inputs and learns to rewrite them . they show that it provides better adversarial robustness than existing defense approaches . |
| Outcome: | The proposed model improves adversarial robustness without compromising task accuracy on a sentiment classification dataset. |
Generative Adversarial Training with Perturbed Token Detection for Model Robustness (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing adversarial training methods use discrete tokens to deceive models . current approaches use embeddings, whereas actual text-based training uses discrete text tokens. |
| Approach: | They propose a framework that integrates gradient-based learning, adversarial example generation and perturbed token detection to enhance adversariarial robustness. |
| Outcome: | The proposed framework surpasses the state-of-the-art results of ChatGPT by 10% in average accuracy. |
On the Relationship between Skill Neurons and Robustness in Prompt Tuning (2024.lrec-main)
Copied to clipboard
| Challenge: | Prompt Tuning is a parameter-efficient finetuning method for pre-trained large language models (PLMs). |
| Approach: | They propose to use RoBERTa to fine tune pre-trained large language models by finetuning only a small set of parameters to adjust for downstream tasks. |
| Outcome: | The proposed method activates specific neurons in the transformer’s feed-forward networks that are highly predictive and selective for the given task. |
Model Unlearning via Sparse Autoencoder Subspace Guided Projections (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing unlearning strategies lack interpretability or fail to provide robust defense against adversarial prompts. |
| Approach: | They propose a framework that leverages SAE features to drive targeted updates in the model’s parameter space. |
| Outcome: | The proposed framework reduces harmful knowledge accuracy by 3.22% compared to baselines and improves adversarial robustness under jailbreak prompts. |
Unveiling Project-Specific Bias in Neural Code Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) based neural code models struggle to generalize effectively to real-world inter-project out-of-distribution data. |
| Approach: | They propose a Cond-Idf measurement to measure the relatedness of a token with a label and its project-specificness. |
| Outcome: | The proposed framework improves both inter-project OOD generalization and adversarial robustness while not sacrificing accuracy on intra-project IID data. |
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs (2026.acl-long)
Copied to clipboard
Dasol Choi, DongGeon Lee, Brigitta Jesica Kartono, Helena Berndt, Taeyoun Kwon, Joonwon Jang, Haon Park, Hwanjo Yu, Minsuk Kahng
| Challenge: | Large language models are being rapidly adopted across a wide range of domains, including healthcare, finance, and the public sector. |
| Approach: | They propose a framework to evaluate whether large language models comply with policies . they apply COMPASS to eight diverse industry scenarios to validate models . |
| Outcome: | The proposed framework evaluates whether LLMs comply with allowlist and denylist policies. |
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Recent studies have highlighted weak-to-strong generalization, where a strong model trained only on a weak model’s labels surpasses the weak model in task performance. |
| Approach: | They propose two fundamental fine-tuning strategies that leverage trustworthiness regularization during the fine-uning of the weak model and the weak-to-strong transfer to improve trustworthy. |
| Outcome: | The proposed models show that they can generalize robustness, fairness, and privacy better when trained on weak models than models trained on strong models. |